Papers with binary classification
French GossipPrompts: Dataset For Prevention of Generating French Gossip Stories By LLMs (2024.eacl-short)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are undergoing a dynamic transformation . however, there is a potential risk of LLMs creating gossips when prompted with contexts . |
| Approach: | a dataset is used to identify prompts that lead to the creation of gossipy content in the french language. |
| Outcome: | a new dataset identifies prompts that lead to the creation of gossipy content in the french language . the model achieves an accuracy of 89.95% . |
Systematic Evaluation of Predictive Fairness (2022.aacl-main)
Copied to clipboard
| Challenge: | Several methods have been proposed to mitigate bias in training on biased datasets. |
| Approach: | They propose to examine the effect of target class imbalance and stereotyping on model performance by analyzing binary classification, profession prediction and regression tasks. |
| Outcome: | The proposed methods show that data conditions have a strong influence on relative model performance. |
“John is 50 years old, can his son be 65?” Evaluating NLP Models’ Understanding of Feasibility (2023.eacl-main)
Copied to clipboard
Himanshu Gupta, Neeraj Varshney, Swaroop Mishra, Kuntal Kumar Pal, Saurabh Arjun Sawant, Kevin Scaria, Siddharth Goyal, Chitta Baral
| Challenge: | Recent work has found that large-scale language models lack commonsense reasoning ability . a dataset evaluating large-level language models is needed to evaluate their understanding of feasibility . |
| Approach: | They propose a question-answering dataset that tests understanding of feasibility . they propose to use commonsense reasoning to reason about when an action is feasible . |
| Outcome: | The proposed dataset shows that state-of-the-art models struggle to answer feasibility questions correctly. |
LLM-DetectAIve: a Tool for Fine-Grained Machine-Generated Text Detection (2024.emnlp-demo)
Copied to clipboard
Mervat Abassy, Kareem Elozeiri, Alexander Aziz, Minh Ta, Raj Tomar, Bimarsha Adhikari, Saad Ahmed, Yuxia Wang, Osama Mohammed Afzal, Zhuohan Xie, Jonibek Mansurov, Ekaterina Artemova, Vladislav Mikhailov, Rui Xing, Jiahui Geng, Hasan Iqbal, Zain Mujahid, Tarek Mahmoud, Akim Tsvigun, Alham Aji, Artem Shelmanov, Nizar Habash, Iryna Gurevych, Preslav Nakov
| Challenge: | a large number of machine-generated texts are often hard to distinguish between human-written and machine-generated text . this raises concerns about potential misuse, especially within educational and academic domains . |
| Approach: | They propose a system that can detect whether a text is human-written or machine-generated . they use a fine-grained classification schema to identify the use of machine-generated text . |
| Outcome: | The proposed system can distinguish between human-written and machine-generated text . it can detect attempts to obfuscate the fact that a text was machine- generated . |
New or Old? Exploring How Pre-Trained Language Models Represent Discourse Entities (2022.coling-1)
Copied to clipboard
| Challenge: | Recent research shows pre-trained language models learn to encode syntactic knowledge to a certain degree. |
| Approach: | They propose to investigate the information-status of entities as discourse-new or discourse-old . they use binary classification and sequence labeling to investigate their ability to encode syntactic knowledge . |
| Outcome: | The proposed models encode information on whether an entity has been introduced before or not in the discourse. |
CVAE-based Re-anchoring for Implicit Discourse Relation Classification (2021.findings-emnlp)
Copied to clipboard
| Challenge: | Existing studies show that training implicit discourse relation classifiers suffers from data sparsity. |
| Approach: | They propose a re-anchoring strategy to reduce the risk of erroneous sampling . they use Conditional VAE to estimate the risk and migrate the anchor to reduce it . |
| Outcome: | The proposed method improves the baseline classifier performance on PDTB v2.0 . |
Advancing Process Verification for Large Language Models via Tree-Based Preference Learning (2024.emnlp-main)
Copied to clipboard
| Challenge: | Existing methods for generating step-by-step rationales fail to fully utilize the relative merits of intermediate steps, limiting the effectiveness of feedback provided. |
| Approach: | They propose a tree-based preference learning verifier that constructs reasoning trees via a best-first search algorithm and collects step-level paired data for preference training. |
| Outcome: | The proposed approach outperforms existing benchmarks on arithmetic and commonsense reasoning tasks. |
Frustratingly Easy System Combination for Grammatical Error Correction (2022.naacl-main)
Copied to clipboard
| Challenge: | Using a simple logistic regression algorithm, we combine GEC models for binary classification. |
| Approach: | They propose a logistic regression algorithm that can combine GEC models with binary classification. |
| Outcome: | The proposed method outperforms the state-of-the-art by 4.2 points on the CoNLL-2014 and 7.2 points on BEA-2019 test sets. |
Detection of Human and Machine-Authored Fake News in Urdu (2025.acl-long)
Copied to clipboard
| Challenge: | Existing methods for fake news detection focus on binary classification and English texts, ignoring the distinction between machine-generated true vs. fake news and low-resource languages. |
| Approach: | They propose to include machine-generated news focusing on Urdu to improve accuracy and robustness. |
| Outcome: | The proposed strategy improves accuracy and robustness across four datasets in various settings. |
Training Classifiers with Natural Language Explanations (P18-1)
Copied to clipboard
| Challenge: | a semantic parser converts explanations into programmatic labeling functions . a standard protocol for obtaining a labeled dataset provides only one bit of information per example . |
| Approach: | They propose a framework where an annotator provides an explanation for each labeling decision . they use a semantic parser to convert these explanations into programmatic labeling functions . |
| Outcome: | The proposed framework trains classifiers faster by providing explanations instead of labels . the proposed framework is based on a rule-based semantic parser . |
Training ELECTRA Augmented with Multi-word Selection (2021.findings-acl)
Copied to clipboard
| Challenge: | Existing pre-training methods for NLP tasks require massive computation resources. |
| Approach: | They propose a method that trains a discriminator to detect replaced tokens and select original tokens from candidate sets. |
| Outcome: | The proposed method improves ELECTRA based on multi-task learning on GLUE and SQUAD datasets. |
M-BRe: Discovering Training Samples for Relation Extraction from Unlabeled Texts with Large Language Models (2025.emnlp-main)
Copied to clipboard
| Challenge: | Existing methods to extract training instances from unlabeled texts are expensive . sentences that contain the target relations in texts can be scarce and difficult to find . |
| Approach: | They propose a framework that can automatically extract training instances from unlabeled texts for RE. |
| Outcome: | The proposed method can extract training instances from unlabeled texts for RE. |
Learning from Natural Language Explanations for Generalizable Entity Matching (2024.emnlp-main)
Copied to clipboard
| Challenge: | Entity matching is the task of linking records from different sources that refer to the same real-world entity. |
| Approach: | They propose to "distill" LLM reasoning into smaller entity matching models via natural language explanations. |
| Outcome: | The proposed model distillation approach achieves strong performance on out-of-domain generalization tests (10.85% F-1). |
Cross-lingual Approaches for the Detection of Adverse Drug Reactions in German from a Patient’s Perspective (2022.lrec-1)
Copied to clipboard
| Challenge: | a recent study shows that the class labels of german documents containing ADRs are imbalanced . clinical trials and physicians prescribing medications cannot cover every potential use case. |
| Approach: | They propose to use binary annotated documents from a german patient forum to detect ADRs. |
| Outcome: | The proposed model achieves an F1 score of 37.52 for the positive class on the German patient forum. |
Hierarchical CVAE for Fine-Grained Hate Speech Classification (D18-1)
Copied to clipboard
| Challenge: | Existing work on automated hate speech detection focuses on binary classification or on differentiating among a small set of categories. |
| Approach: | They propose a method to discriminate among 40 hate groups of 13 different hate group categories. |
| Outcome: | The proposed method outperforms discriminative models on a fine-grained hate speech classification task. |
LGAR: Zero-Shot LLM-Guided Neural Ranking for Abstract Screening in Systematic Literature Reviews (2025.findings-acl)
Copied to clipboard
| Challenge: | Existing methods for abstract screening focus on binary classification settings; existing question answering (QA) based ranking approaches suffer from error propagation. |
| Approach: | They propose a systematic literature review (SLR) method that uses large language models to evaluate the SLR's inclusion and exclusion criteria. |
| Outcome: | The proposed method outperforms existing question answering (QA) based methods by 5-10 pp. in mean precision. |
A Multi-Task Approach for Improving Biomedical Named Entity Recognition by Incorporating Multi-Granularity information (2021.findings-acl)
Copied to clipboard
| Challenge: | Neural named entity recognition (BioNER) methods require large amount of annotated data, while the annotating BioNER datasets are often difficult to obtain and small in scale due to the limitations of privacy, ethics and high degree of specialization. |
| Approach: | They propose a method that utilizes latent multi-granularity information in annotated bioNER datasets to alleviate the lack of training samples. |
| Outcome: | The proposed model improves over the BioBERT baseline and can get more than 3% improvement of F1score in low-resource scenarios. |
Misogyny and Aggressiveness Tend to Come Together and Together We Address Them (2022.lrec-1)
Copied to clipboard
| Challenge: | Using a binary task to identify whether a tweet is misogynous and aggressive, we compare two approaches to address these problems: one multi-class model that discriminates between all the classes at once; and a cascaded approach where the binary classification is carried out separately. |
| Approach: | They propose a multi-class model that discriminates between all the classes at once and a cascaded approach where the binary classification is carried out separately and then joined together. |
| Outcome: | The proposed models outperform the top submissions to Evalita on the 2020 shared task on automatic misogyny and aggressiveness identification in Italian tweets. |
Improving Low-Resource Named Entity Recognition using Joint Sentence and Token Labeling (2020.acl-main)
Copied to clipboard
| Challenge: | Existing models for named entity recognition (NER) use sentence-level labels, which are expensive to obtain, to improve NER. |
| Approach: | They propose a sentence-level named entity recognition model that uses sentence-based labels that are easy to obtain. |
| Outcome: | The proposed model produces 3.78%, 4.20%, 2.08% improvements in F1 over the baseline on e-commerce product titles in Vietnamese, Thai, and Indonesian, respectively. |
Time-RA: Towards Time Series Reasoning for Anomaly Diagnosis with LLM Feedback (2026.findings-acl)
Copied to clipboard
Yiyuan Yang, Zichuan Liu, Lei Song, Kai Ying, Stephen Wang, Joshua Thomas Bamford, Svitlana Vyetrenko, Jiang Bian, Qingsong Wen
| Challenge: | Time series anomaly detection (TSAD) has traditionally focused on binary classification and lacks the fine-grained categorization and explanatory reasoning required for transparent decision-making. |
| Approach: | They propose a time-series reasoning task that reformulates TSAD from discriminative to reasoning-intensive paradigm. |
| Outcome: | The proposed task reformulates TSAD from discriminative to reasoning-intensive paradigm. |
Among Us: Language of Conspiracy Theorists on Mainstream Reddit (2026.acl-long)
Copied to clipboard
| Challenge: | Conspiracy theories are influential, alternative narratives that explain events through the actions of secretive, malevolent groups. |
| Approach: | They analyze a large-scale longitudinal dataset of over 500 million comments on reddit . they show that users exhibit distinctive linguistic patterns that enable machine learning models to distinguish them from the general population within individual communities. |
| Outcome: | The proposed model outperforms global classifiers by 17 percentage points. |
Detection of Reading Absorption in User-Generated Book Reviews: Resources Creation and Evaluation (2020.lrec-1)
Copied to clipboard
Piroska Lendvai, Sándor Darányi, Christian Geng, Moniek Kuijpers, Oier Lopez de Lacalle, Jean-Christophe Mensonides, Simone Rebora, Uwe Reichel
| Challenge: | a new study aims to detect how and when readers are experiencing engagement with a literary work . empirical literary studies and language technology are used to investigate reading absorption . |
| Approach: | They annotated user-generated book reviews with reading absorption categories . they then performed supervised binary classification of the mental state of absorption . |
| Outcome: | The proposed corpus of user-generated reviews is compared with machine learning models and a benchmark corpus. |
Exploring BERT-Based Classification Models for Detecting Phobia Subtypes: A Novel Tweet Dataset and Comparative Analysis (2024.lrec-main)
Copied to clipboard
| Challenge: | Phobias are characterized by an intense and irrational fear of specific objects, situations, or activities despite there being no real risk or only a minor threat involved. |
| Approach: | They propose to use a dataset of 811,569 English tweets from user timelines spanning 102 phobia subtypes over six months to classify users into 65 specific phobias. |
| Outcome: | The proposed dataset includes 47,614 self-diagnosed phobia users and a high f1 score for binary classification and multi-class classification. |
Exploring Pathological Speech Quality Assessment with ASR-Powered Wav2Vec2 in Data-Scarce Context (2024.lrec-main)
Copied to clipboard
| Challenge: | Current studies only gain good results on simple tasks such as binary classification due to data scarcity. |
| Approach: | They propose to use the pre-trained Wav2Vec2 architecture for both SSL, and ASR as feature extractor in speech assessment. |
| Outcome: | The proposed system achieves the best results on the HNC dataset using 95 training samples. |
SharedCon: Implicit Hate Speech Detection using Shared Semantics (2024.findings-acl)
Copied to clipboard
| Challenge: | Recent studies suggest that classifying hateful posts in a binary manner may not address nuanced task of detecting implicit hate speech. |
| Approach: | They propose a contrastive learning approach that leverages shared semantics among data to detect implicit hate speech. |
| Outcome: | The proposed approach is based on a clustering-based contrastive learning approach with human-written implications or machine-generated augmented data. |
A Tale of Evaluating Factual Consistency: Case Study on Long Document Summarization Evaluation (2025.findings-acl)
Copied to clipboard
| Challenge: | Despite the recent progress for summarization models in producing fluent summaries, they still encounter challenges when long sequences of generated texts and inputs (over thousands of words) need to be evaluated. |
| Approach: | They conduct a systematic analysis of factual-consistency evaluation systems across four long-document datasets and examine the relationship between sentence-level and summary-level model performance. |
| Outcome: | The proposed models can achieve higher recall in error detection for older summaries, yet struggle with false positives and fine-grained error detection. |
Building a Sentiment Corpus of Tweets in Brazilian Portuguese (L18-1)
Copied to clipboard
| Challenge: | Sentiment analysis is a popular area of Natural Language Processing due to its subjective and semantic characteristics. |
| Approach: | They propose to annotate Brazilian Portuguese sentences manually using a sentiment corpus . they run experiments on polarity classification using six machine learning classifiers . |
| Outcome: | The proposed method is based on a Brazilian Portuguese sentiment corpus and achieved 80.38% on F-Measure and 64.87% when including the neutral class. |
Tackling Irony Detection using Ensemble Classifiers (2022.lrec-1)
Copied to clipboard
| Challenge: | Automated approaches to irony detection still fall short of what one would consider desirable performance. |
| Approach: | They propose to use transformer-based approaches to automate irony detection in social media . they propose to augmentation training data to address the binary and fine-grained problem . |
| Outcome: | The proposed methods improve performance over baselines and are not decisive for good results. |
Debate-to-Detect: Reformulating Misinformation Detection as a Real-World Debate with Large Language Models (2025.emnlp-main)
Copied to clipboard
| Challenge: | Despite advances in large language models, their application to misinformation detection remains hindered by issues of logical inconsistency and superficial verification. |
| Approach: | They propose a multi-agent debate framework that reformulates misinformation detection as a structured adversarial debate based on fact-checking workflows . |
| Outcome: | The proposed framework enables iterative refinement of evidence while improving decision transparency. |
HateBR: A Large Expert Annotated Corpus of Brazilian Instagram Comments for Offensive Language and Hate Speech Detection (2022.lrec-1)
Copied to clipboard
| Challenge: | In Brazil, hate speech is prohibited, however the regulation is not effective due to the difficulty of identifying, quantifying and classifying this kind of online content. |
| Approach: | They propose to annotate a large corpus of Brazilian Instagram comments manually and to use it to detect hate speech and offensive language. |
| Outcome: | The HateBR corpus was collected from the comment section of Brazilian politicians’ accounts on Instagram and manually annotated by specialists, reaching a high inter-annotator agreement. |
Explaining Matters: Leveraging Definitions and Semantic Expansion for Sexism Detection (2025.acl-long)
Copied to clipboard
| Challenge: | Existing tools for sexism detection fail to capture subtle distinctions within sexist content, limiting their practical applicability. |
| Approach: | They propose two techniques to address class imbalance and nuanced nature of sexist language . definition-based data augmentation leverages category-specific definitions to generate semantically-aligned examples . |
| Outcome: | The proposed techniques improve accuracy across all tasks and improve reliability. |
More Than Sum of Its Parts: Deciphering Intent Shifts in Multimodal Hate Speech Detection (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing systems struggle with multimodal content where the emergent meaning transcends the aggregation of individual modalities. |
| Approach: | They propose a framework to characterize semantic intent shifts where modalities interact to construct implicit hate from benign cues or neutralize toxicity through semantic inversion. |
| Outcome: | The proposed framework outperforms state-of-the-art benchmarks on H-VLI and on established benchmarks. |
Word-Level Detection of Code-Mixed Hate Speech with Multilingual Domain Transfer (2025.findings-acl)
Copied to clipboard
| Challenge: | a growing problem in language detection tasks is code-mixing, a combination of more than one language . lack of available datasets for code-mixing causes the problem . authors propose a multilingual approach to code-matching . |
| Approach: | They propose to use an annotated hate speech dataset to detect code-mixing in profane language . they propose to apply bilingual fine-tuned models to code-mixed hate speech in german rap lyrics . |
| Outcome: | The proposed model can detect code-mixed hate speech and neologisms in German rap lyrics . the proposed model is more nuanced than binary classification . |
Generating Attribution Reports for Manipulated Facial Images: A Dataset and Baseline (2026.acl-long)
Copied to clipboard
| Challenge: | Existing facial forgery detection methods focus on binary classification or pixel-level localization, providing little semantic insight into the nature of the manipulation. |
| Approach: | They propose a multimodal task that localizes forged regions and generates natural language explanations grounded in editing process. |
| Outcome: | The proposed task localizes forged regions and generates natural language explanations grounded in editing process. |
Authorship Attribution in Multilingual Machine-Generated Texts (2026.acl-long)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have reached human-like fluency and coherence, but distinguishing machine-generated text from human-written content becomes increasingly difficult. |
| Approach: | They propose a problem of multilingual authorship attribution (AA) that involves attributing texts to human or multiple LLM generators across diverse languages. |
| Outcome: | The proposed method can be adapted to multilingual settings, but still has significant limitations and challenges. |